Repository navigation
fix(codex): keep system prompts in input for GPT-5 automatic prompt caching - #1346
Merged
diegosouzapw merged 1 commit intoApr 16, 2026
Merged
Conversation
OpenAI's automatic prompt caching for GPT-5 models computes the cache
key from the serialized `input` array and `tools`, but does NOT include
the `instructions` field. The current code moves system messages from
`input` into `instructions` via `hoistSystemMessagesToInstructions()`,
which removes them from the cacheable prefix entirely.
This causes 0% cache hit rates for native Responses API passthrough
clients that send system prompts as role=system messages in `input`.
Changes:
- Add `convertSystemToDeveloperRole()` — converts system → developer
role in-place within `input` (Codex accepts developer but rejects
system). This keeps the content in the cacheable prefix.
- For native passthrough requests: use the new function instead of
hoisting to `instructions`. Set a minimal placeholder instruction
instead of injecting CODEX_DEFAULT_INSTRUCTIONS.
- For translated requests (Chat Completions → Responses): preserve
the existing hoist + default instructions behavior (no change).
Before: instructions="3000-word default + system prompt", input=[user msgs only]
→ cached_tokens = 0 (instructions not in cache key)
After: instructions="minimal", input=[{role:"developer",...}, user msgs...]
→ cached_tokens = system_prompt + tools + conversation prefix
Ref: https://community.openai.com/t/caching-is-borked-for-gpt-5-models/1359574
Ref: https://community.openai.com/t/no-caching-with-model-responses/1338627
Contributor
There was a problem hiding this comment.
Code Review
This pull request implements a cache-aware strategy for system prompt handling in the CodexExecutor to optimize for GPT-5 models. It introduces the convertSystemToDeveloperRole function, which converts system roles to developer roles within the input array during native passthrough to preserve OpenAI's prompt caching. For translated requests, the existing behavior of hoisting system messages to the instructions field is maintained. I have no feedback to provide.
Owner
|
Thanks @Gi99lin for this great contribution! 🎉 This PR has been evaluated and successfully integrated into the release/v3.6.7 branch and will be part of the final production release. We appreciate your effort! |
This was referenced Apr 17, 2026
Merged
Poid-ZA
pushed a commit
to Poid-ZA/OmniRoute
that referenced
this pull request
Aug 5, 2026
…egosouzapw#1346) Integrated into release/v3.6.7
muhamadgalihsaputra
pushed a commit
to niyatna/NiyatnaRoute
that referenced
this pull request
Sep 27, 2026
…egosouzapw#1346) Integrated into release/v3.6.7
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
OpenAI's automatic prompt caching for GPT-5 models returns
cached_tokens: 0on every request routed through OmniRoute, even with identical 53K-token prompts sent seconds apart.Root cause: The
instructionsfield in the Responses API is not included in the prompt cache key computation for GPT-5 models. OpenAI only caches based on the serializedinputarray +tools. The current code moves system messages frominput(cacheable) intoinstructions(not cacheable) viahoistSystemMessagesToInstructions(), destroying the cache prefix entirely.Evidence
Diagnostic logging on a production deployment confirmed:
Three consecutive requests with identical
instructions + toolsprefix, same model, within 90 seconds — zero cache hits. Meanwhile, a MiniMax provider using the same proxy achieved cache hits immediately:This is consistent with community reports:
What this PR changes
For native Responses API passthrough requests (
_nativeCodexPassthrough === true):inputtoinstructionssystem→developerrole, kept ininputinstructionsfieldCODEX_DEFAULT_INSTRUCTIONS(3000+ words)toolsonly (system prompt removed frominput)developer msg + tools + conversation prefixFor translated requests (Chat Completions → Responses format): no change. The existing hoist + default instructions behavior is preserved.
New function:
convertSystemToDeveloperRole()Converts
role: "system"→role: "developer"in-place within theinputarray. This is necessary because:systemrole ininput(existing comment in code confirms this)developerrole as the replacementinputpreserves the cacheable prefixImpact
For a deployment routing ~7K daily Codex requests averaging 41K input tokens each:
Testing
To verify the fix:
cached_tokensinresponse.completedSSE eventscached_tokens > 0Example diagnostic (add to proxy/middleware):